Papers with statistical framework
Deep Neural Networks at the Service of Multilingual Parallel Sentence Extraction (C18-1)
Copied to clipboard
| Challenge: | Existing models for parallel data harvesting from Wikipedia are language-independent, robust and highly scalable. |
| Approach: | They propose an end-to-end neural model for large-scale parallel data harvesting from Wikipedia . their model is language-independent, robust, and highly scalable . |
| Outcome: | The proposed model is language-independent, robust, and highly scalable. |
Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation (2024.emnlp-main)
Copied to clipboard
Zhuohang Li, Jiaxin Zhang, Chao Yan, Kamalika Das, Sricharan Kumar, Murat Kantarcioglu, Bradley Malin
| Challenge: | Language models suffer from poor interpretability and transparency, as well as the intrinsic risk of hallucination and misinformation. |
| Approach: | They propose a statistical framework that assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge. |
| Outcome: | The proposed framework assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge. |
Verifiable LLM-Generated Text Detection via Projected Semantic-Structural Distributions (2026.acl-long)
Copied to clipboard
Ruochong Xiong, Qien Li, Wangwang Lian, Yulong Wan, Hanlin Xue, Zhouxing Tan, Han Yang, Fengyu Lu, Junfei Liu
| Challenge: | Existing methods for detecting LLM-Generated text suffer from distribution misalignment and limited interpretability. |
| Approach: | They propose a statistical framework utilizing supervised subspace learning to extract compact features and construct conditional semantic distributions based on syntactic structures. |
| Outcome: | The proposed framework is superior in cross-domain, cross-model, and adversarial scenarios. |
Analyzing and Modeling LLM Response Lengths with Extreme Value Theory: Anchoring Effects and Hybrid Distributions (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches treat length as an incidental output property rather than a statistically regular phenomenon worthy of rigorous modeling. |
| Approach: | They propose a statistical framework for modeling and controlling large language model response lengths using extreme value theory and cross-validation on Qwen and DeepSeek architectures. |
| Outcome: | The proposed model improves tail fit and generalizability while maintaining generalizzability. |